audit: deploys_audit table + admin endpoint — answers what was running when - #57
Merged
Merged
Conversation
…g when Adds an append-only deploy-identity log so the founder can answer "what image was serving traffic at $TIME on service $X?" — a question /healthz can't answer once a pod has rolled. One row per unique (service, commit_id, image_digest) tuple, written by the binary itself on startup via ON CONFLICT DO NOTHING, served at GET /api/v1/<admin- prefix>/deploys behind the existing RequireAdmin + ADMIN_PATH_PREFIX gates. Migration 022_deploys_audit.sql creates the table + unique index (backs the self-report INSERT's ON CONFLICT) + service+time index (backs the admin endpoint's default sort). The image digest source is the IMAGE_DIGEST env var, populated by k8s via valueFrom.fieldRef on status.containerStatuses[0].imageID — unset (local dev) falls back to the literal "local-build" sentinel so dev boots collapse onto one row instead of being randomly attributed. The /healthz handler is left alone — INSERTs on every kube probe would hammer the platform DB. Self-report fires exactly once per process at startup; the unique index defends against accidental re-fires. Tests: 9 new model tests pin the SQL contract (basic insert, dedup on same tuple, two rows for different digests, buildinfo sentinels → NULL, identity-field validation, service filter, ORDER BY DESC, since filter, invalid service rejected). 8 new handler tests pin the HTTP contract (admin-closed-by-default, non-admin → 403, empty table → [], single-row round-trip, service filter, invalid service/since/limit → 400). 3 new main-package tests pin the IMAGE_DIGEST resolver (unset → local-build, empty → local-build, real value passes through). All 20 new tests green; `make test-unit` all packages green. Worker + provisioner self-reports are out of scope for this PR — those services live in sibling repos (InstaNode-dev/worker, InstaNode-dev/provisioner) and own their own startup paths. The model exposes DeployServiceWorker / DeployServiceProvisioner constants so their startup hooks can call InsertSelfReport against the shared platform DB once they pull this migration through their own RunMigrations pipelines. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
5 of 6 tasks
mastermanas805
added a commit
that referenced
this pull request
May 14, 2026
…obby_plus copy (#107) Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer- facing restore stops being one accidental retry away from a destroyed database. #57/#Q45 — restore replay guard Adds models.HasInflightRestore; a second POST /restore for the same resource while a prior row is pending/running returns 409 restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB error (concurrent pg_restore --clean races itself). #58/#A2 — restore to a new DB Accepts optional target_resource_id. Worker restores into the target (same team only); audit row carries both source + target ids. Skips the destructive-ack ceremony because the agent already opted into a fresh database. #59 — backup integrity Migration 043 adds a nullable sha256 TEXT column to resource_backups. Worker stamps the digest at finalize; restore handler verifies before pg_restore. NULL on legacy rows is logged + accepted (fail-open on pre-043 data). #64/#Q46 — cross-tenant 404 GetBackupByIDForTeam joins resources to scope by team; cross-tenant backup_id guess now returns 404 backup_not_found instead of 400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture. #65/#Q47 — refund quota on failure New internal endpoint POST /internal/teams/:id/backup-quota/refund (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The worker calls this when a MANUAL backup fails terminally so the team's daily manual-backups counter is credited back. #66/#Q48 — hobby agent_action points to Hobby Plus AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402 now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping the customer past the cheapest restore-enabled plan onto Pro ($49). #67/#Q49 — destructive ack required for in-place restore In-place restore (no target_resource_id) now requires destructive_acknowledgment: true in the body. pg_restore --clean drops every table — refusing without an explicit ack prevents an agent testing a backup from wiping a live customer DB. #Q50 — RPO/RTO on /capabilities Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes (added in instant.dev/common/plans#13). Tests - +5 backup_test.go cases: ReplayBlocked, TargetNewDB, RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus, CrossTenantBackupID_404 - All prior backup tests updated to include destructive_acknowledgment - Contract test enforces the four new agent_action constants DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
mastermanas805
added a commit
that referenced
this pull request
May 21, 2026
…obby_plus copy Wave FIX-H wraps up the BugBash B36 backup-correctness items so customer- facing restore stops being one accidental retry away from a destroyed database. #57/#Q45 — restore replay guard Adds models.HasInflightRestore; a second POST /restore for the same resource while a prior row is pending/running returns 409 restore_in_progress + AgentActionRestoreInflight. Fail-CLOSED on DB error (concurrent pg_restore --clean races itself). #58/#A2 — restore to a new DB Accepts optional target_resource_id. Worker restores into the target (same team only); audit row carries both source + target ids. Skips the destructive-ack ceremony because the agent already opted into a fresh database. #59 — backup integrity Migration 043 adds a nullable sha256 TEXT column to resource_backups. Worker stamps the digest at finalize; restore handler verifies before pg_restore. NULL on legacy rows is logged + accepted (fail-open on pre-043 data). #64/#Q46 — cross-tenant 404 GetBackupByIDForTeam joins resources to scope by team; cross-tenant backup_id guess now returns 404 backup_not_found instead of 400 backup_resource_mismatch. Matches FIX-B tenant-isolation posture. #65/#Q47 — refund quota on failure New internal endpoint POST /internal/teams/:id/backup-quota/refund (WORKER_INTERNAL_JWT_SECRET HS256, fail-closed when unset). The worker calls this when a MANUAL backup fails terminally so the team's daily manual-backups counter is credited back. #66/#Q48 — hobby agent_action points to Hobby Plus AgentActionRestoreRequiresHobbyPlus added. The Hobby-tier restore 402 now nudges Hobby Plus ($19/mo, restore enabled) instead of skipping the customer past the cheapest restore-enabled plan onto Pro ($49). #67/#Q49 — destructive ack required for in-place restore In-place restore (no target_resource_id) now requires destructive_acknowledgment: true in the body. pg_restore --clean drops every table — refusing without an explicit ack prevents an agent testing a backup from wiping a live customer DB. #Q50 — RPO/RTO on /capabilities Adds rpo_minutes + rto_minutes per tier in plans.yaml (anonymous/free = 0/0, hobby/hobby_plus = 1440/30, pro/team = 60/15) wired into the /api/v1/capabilities matrix via plans.Registry.RPOMinutes/RTOMinutes (added in instant.dev/common/plans#13). Tests - +5 backup_test.go cases: ReplayBlocked, TargetNewDB, RequiresDestructiveAck, HobbyAgentActionPointsToHobbyPlus, CrossTenantBackupID_404 - All prior backup tests updated to include destructive_acknowledgment - Contract test enforces the four new agent_action constants DO NOT TOUCH list respected — no edits to email/, dpop.go, circuit/. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds an append-only deploy-identity log so the founder can answer
"what image was serving traffic at $TIME on service $X?" — a
question
/healthzcan't answer once a pod has rolled.deploys_audittable with a unique index on(service, commit_id, image_digest)backing the self-reportON CONFLICT DO NOTHINGand a(service, applied_at DESC)indexbacking the read path.
main.gocallsmodels.InsertSelfReportexactlyonce at process startup. Idempotent: same image on N replicas =
one row.
GET /api/v1/<admin-prefix>/deploys?service=&since=&limit=registered behind the existing
RequireAdmin+ADMIN_PATH_PREFIXgates. Intentionally omitted from
/openapi.json(admin-obscuritypattern, see
internal/handlers/openapi.go).IMAGE_DIGESTenv var, populated by k8svia
valueFrom.fieldRef: status.containerStatuses[0].imageID.Unset (local dev) →
"local-build"sentinel.Iron rules
make test-unitgreen./healthzdoes NOT do the INSERT — startup-only.Test plan
testhelpers.runMigrationsmirrors it)ON CONFLICTprevents duplicateapplied_at DESCagent_action?service=filter works (and rejects unknown values with 400)IMAGE_DIGESTunset → row uses"local-build""dev","unknown") → NULL columnssince/limit→ 400Worker + provisioner
Out of scope for this PR — those services live in sibling repos
(
InstaNode-dev/worker,InstaNode-dev/provisioner). The modelexports
DeployServiceWorker/DeployServiceProvisionerconstantsso their startup hooks can call
models.InsertSelfReportagainst theshared platform DB once they pull this migration through their own
RunMigrationspipelines.🤖 Generated with Claude Code